The Lixto Project: Exploring New Frontiers of Web Data Extraction

نویسندگان

  • Julien Carme
  • Michal Ceresna
  • Oliver Frölich
  • Georg Gottlob
  • Tamir Hassan
  • Marcus Herzog
  • Wolfgang Holzinger
  • Bernhard Krüpl
چکیده

The Lixto project is an ongoing research effort in the area of Web data extraction. Whereas the project originally started out with the idea to develop a logic-based extraction language and a tool to visually define extraction programs from sample Web pages, the scope of the project has been extended over time. Today, new issues such as employing learning algorithms for the definition of extraction programs, automatically extracting data from Web pages featuring a table-centric visual appearance, and extracting from alternative document formats such as PDF are being investigated.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

I-39: Exploring New Frontiers in Human Y Chromosome Proteome Project

The major goal of the Chromosome-Centric Human Proteome Project (C-HPP) is to systematically map the entire human proteome with the intent to enhance our understanding of human biology at the cellular level. However, this goal may be hindered by the lack of quality observations of given proteins due to absence of expression in a given tissue, very low abundance, and expression only in rare samp...

متن کامل

Visual Web Information Extraction with Lixto

We present new techniques for supervised wrapper generation and automated web information extraction, and a system called Lixto implementing these techniques. Our system can generate wrappers which translate relevant pieces of HTML pages into XML. Lixto, of which a working prototype has been implemented, assists the user to semi-automatically create wrapper programs by providing a fully visual ...

متن کامل

Logic, Languages, and Rules for Web Data Extraction and Reasoning over Data

This paper gives a short overview of specific logical approaches to data extraction, data management, and reasoning about data. In particular, we survey theoretical results and formalisms that have been obtained and used in the context of the Lixto Project at TU Wien, the DIADEM project at the University of Oxford, and the VADA project, which is currently being carried out jointly by the univer...

متن کامل

Web Information Acquisition with Lixto Suite: A Demonstration∗

We demonstrate the Lixto Suite, a web data extraction and transformation software kit for retrieving and converting information from various sources to various customer devices. With the Lixto Suite, non-technical content managers can rapidly develop applications in the areas of M-Commerce, E-Commerce, content integration and corporate portals.

متن کامل

Intelligent Wrapping from PDF Documents

Wrapping is the process of navigating a data source, semiautomatically extracting data and transforming it into a form suitable for data processing applications. The semi-structured form of web pages, coupled with the availability of business-relevant data, has led to the availability of several established products on the market for wrapping data from the Web. One such approach is the Lixto me...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2006